Octopus · by Mister.D

Your delivery pipeline, run by agents. Governed by you.

Every task moves through analysis, development, testing and review with a Claude agent on each stage, an auditor that can block the work, and a person on your team who signs off before anything merges.

Claude-powered·Isolated sandboxes·Runs on your servers

  • 1 command to self-host: octopus init
  • 24h max lifetime of each run’s scoped API key
  • up to 450% faster QA with annotated video evidence

Workflow automation

Your delivery process, running as a pipeline.

Define the stages your team already uses. Octopus puts a specialized agent on each one, an auditor on every handoff and a human at the gates that matter. A ticket goes in; a reviewed, tested, deployed change comes out.

  1. 01 · Intake

    A ticket lands on the board

    From the board, the API or an automation. It enters the first stage of a pipeline you configured, with its own agents and rules.

  2. 02 · Analysis

    The analysis agent reads ticket and repo

    It checks the code the ticket touches, writes acceptance criteria and turns them into concrete test cases before anyone writes a line.

  3. 03 · Planning

    A plan the auditor can check

    Files to change, subtasks, risks. The auditor agent approves it, asks for changes or blocks it with a reason written on the task.

  4. 04 · Development

    Code in an isolated sandbox

    Each run gets its own container or microVM and a short-lived API key scoped to that run. The agent opens a branch, commits and opens the pull request.

  5. 05 · Testing

    Test cases run in a real browser

    The testing agent executes every case, captures screenshots and records a video annotated with each step and result. All of it lands on the task.

  6. 06 · Review & deploy

    Auditor, then a human, then production

    The auditor reviews the PR, a person approves at the review gate and the merge triggers the deploy through GitHub Actions.

Automations

When something happens, the next step runs itself.

Triggers watch the board: new task, stage change, comment, PR merged or a schedule. Conditions narrow them. Actions assign an agent, move a stage, notify or run a pipeline.

automation.yaml

QA on entry

When Task moves to Test
If Board = Web app
Then
Run QA agentAttach video to taskNotify reviewer

Ship on merge

When Pull request merged
If Branch = main
Then
Deploy staging via GitHub ActionsMove task to DoneNotify channel

Weekly report

When Every Monday 9:00
If Workspace = Product
Then
Run metrics agentBuild dashboardComment summary
  • 01

    Custom stages per board

    Name them, order them, add or remove them. Each board runs its own process.

  • 02

    One agent per stage

    Each stage has its own prompt, tools and MCP permissions, down to read or write per tool.

  • 03

    Human approval gates

    Pick the stages where a person signs off. Nothing merges past a gate without it.

  • 04

    Triggers and schedules

    Events or cron start the work. Nobody has to remember to push the ticket forward.

Testing & QA

Every test case comes back as an annotated video.

Octopus reads the ticket and its acceptance criteria, writes the test cases, runs each one in a real browser and attaches the recording to the task. Each step is marked on the video: what was clicked, what was typed, what was asserted, and where it broke. QA watches, then approves or comments.

up to 450% faster QA reviews 45 min of manual reproduction per ticket, down to ~10 min watching the video and signing off.
  1. 01 Analyze

    Test cases from the ticket

    The testing agent reads the ticket and the acceptance criteria, then writes the cases, edge cases included, each with a priority.

    CHK-142 · Discount codes + declined cards
    • TC-01 Login with valid credentials P1
    • TC-02 Add item to cart P1
    • TC-03 Apply discount code P0
    • TC-04 Pay with Safari 17 P1
    • TC-05 Empty cart edge case P2
    • TC-06 Error message on declined card P0
  2. 02 Execute

    Run in a real browser

    Every case runs in a real browser inside the run’s isolated sandbox. Clicks, typing, navigation, console and network are all recorded.

    Running 6 cases
    • TC-01 0:12 ✓
    • TC-02 0:21 ✓
    • TC-03 0:34 ✓
    • TC-04 0:29 ✓
    • TC-05 0:09 ✓
    • TC-06 0:48 ✗
  3. 03 Deliver

    Video + report on the task

    An annotated video per case and a pass/fail report land on the task. Failures open a comment with the exact step that broke.

    Attached to CHK-142
    tc-06_declined-card.webm 0:48 · 7 annotations
    test-report.pdf 5 ✓ 1 ✗ · 5 passed · 1 failed
    TC-06 failed at step 4: no “Card declined” banner after payment.
CHK-142 · checkout run · annotated REC
00:00 / 00:16
assertion passed assertion failed issue flagged

Where the 450% comes from

Manual QA 45 min
  • Read ticket
  • Set up data
  • Reproduce
  • Screenshot
  • Write report
With Octopus ~10 min
  • Watch 2-min annotated video
  • Approve or comment

4.5× faster per ticket

What QA gets from every run

  • Annotated video per test case
  • Screenshot at every assertion
  • Pass/fail report
  • Console and network errors captured
  • Reproducible steps, in order
  • All of it attached to the task
Book a demo We run it on one of your real tickets.

Features

Everything the pipeline needs, in one place

Sixteen pieces that take a task from the board to production, with an agent on every stage and a person on every gate.

01 · BOARDS

Boards with your own stages

Name the columns after how your team ships. Each stage decides which agent runs and who signs off.

02 · AGENTS

One specialist per stage

Analysis, planning, development, testing and review each get their own prompt, tools and permissions.

03 · AUDITOR

An auditor on every hand-off

A second agent reads the output and answers approve, request changes or block.

04 · GATES

Humans approve the gates

Nothing reaches review or production until a person on the team clicks approve.

05 · GITHUB APP

Branches, commits, PRs and deploys

Agents open a branch, commit, open the pull request and trigger your GitHub Actions deploy through the Octopus GitHub App.

06 · BROWSER

A real browser with a camera

Agents navigate your app, take screenshots and record video with annotations drawn on the exact element that broke.

07 · ISOLATION

A fresh sandbox per run

Every run gets its own container or microVM on Kubernetes, then it is torn down.

08 · MCP

Tools with per-tool permissions

GitHub, browser, code sandbox and the Octopus API over MCP. Grant read or write tool by tool.

09 · AUTOMATIONS

Triggers that move the work

New task, stage change, comment, merged PR or a schedule fires: assign an agent, move a stage, notify, run the pipeline.

10 · TASK CHAT

Talk to the agent inside the task

Ask why, give context or redirect it without leaving the ticket.

11 · EVIDENCE

Proof attached to every task

Screenshots, PDFs, annotated videos and logs land on the task that produced them.

12 · METRICS

Custom metrics and dashboards

Define the numbers you care about, such as lead time, rework or approvals per stage, and chart them per workspace.

13 · BYOK

Your own Claude key

Agents run on Claude with your Anthropic key. Your usage, your bill, your limits.

14 · WORKSPACES

Workspaces, roles, scoped keys

API keys belong to one workspace with explicit permissions. Each run gets a short-lived key of its own.

15 · IDENTITY

SSO, passkeys and an audit log

SAML single sign-on and passkeys through MisterD ID. Every sensitive action is written to the audit log.

16 · DEPLOY

Runs on your servers

Installs on your infrastructure with Docker Compose or Helm. Team features through a license key. Export and import a whole workspace.

Governance & security

Agents do the work.
You keep the keys.

Every run starts in a fresh sandbox, gets a key scoped to one task, touches only the tools you allowed, and passes an auditor and a human before anything ships. Every step lands in the audit log.

Anatomy of a runrun#8812
01 · SANDBOX microVM · isolated · wiped after the run 24h 02 · SCOPED API KEY scope=task:1284 expires in 24h 03 · ALLOWED MCP TOOLS GitHub @write Browser @use Sandbox @exec Octopus API @read 04 · AUDITOR AGENT approve changes block 05 · HUMAN GATE you approve the merge Approve
audit.log● live
  • 14:02:11 run#8812 key minted scope=task:1284 ttl=24h
  • 14:02:12 run#8812 sandbox up microvm=kata-7f3a
  • 14:02:19 run#8812 mcp github.branch_create @write ok
  • 14:04:37 run#8812 mcp browser.screenshot ok
  • 14:06:02 run#8812 mcp octopus.task_get @read ok
  • 14:06:03 run#8812 mcp octopus.task_update denied scope=@read
  • 14:09:48 run#8812 mcp github.pull_request_open @write ok
  • 14:10:05 auditor verdict=approve stage=review
  • 14:12:40 human gate approved by reviewer role=admin
  • 14:12:41 run#8812 key revoked sandbox destroyed

Six guarantees on every run

  • Isolation per run

    Each run gets its own container or microVM on Kubernetes. Nothing carries over to the next one.

  • Short-lived scoped keys

    A fresh API key per run, limited to its task and workspace, expiring on its own.

  • Per-tool permissions

    Grant GitHub write, browser, sandbox or Octopus API read tool by tool. Read and write scopes are separate.

  • Auditor + human gates

    An auditor agent approves, requests changes or blocks each stage. People sign off at review gates.

  • Full audit log

    Keys minted, tools called, verdicts and approvals, with timestamps and the run that did it.

  • SSO (SAML) & passkeys

    Sign in through MisterD ID. Workspaces with roles and workspace-scoped API keys.

Installed on your servers. Licensed for your team.

Your infrastructure

Everything runs on your servers

Octopus installs where you decide: your cloud account, your data center or a single server. Code, sandboxes, evidence and logs stay there. The only outbound call goes to Claude, with your own key.

~/octopus
$ octopus init✓ pulled images (daemon, web, postgres)✓ workspace created, admin key minted→ Octopus ready at http://localhost:8585
  • Docker Compose in one command, or Kubernetes with Helm
  • Your Claude key: model usage is billed to your account
  • Export and import a whole workspace
Teams

Licensed per team

A license key turns on the team features in your own installation. Pricing follows team size; we size it with you in the demo.

  • Multiple workspaces with roles
  • SSO providers and audit log export (CSV, JSON, SIEM)
  • Helm chart for Kubernetes and priority support
Book a demo

Licensing

Your servers. Your Claude key.
One license per team.

You install Octopus on your own servers, add a team license key to turn on the team features, and connect your own Claude API key. Code, sandboxes, evidence and logs stay with you. The only outbound call goes to Claude.

  1. 01

    Install on your servers

    Your cloud account, your data center or a single server. One command with Docker Compose, or a Helm chart on Kubernetes. Code, sandboxes, evidence and logs stay there.

  2. 02

    Activate your team license

    Install the license key on your own installation and the team features turn on. The key has an expiry date; to renew, you install a new key. No reinstall.

  3. 03

    Connect your Claude key

    Octopus calls Claude with your own API key. Anthropic bills model usage straight to your account. We add no markup.

Base install vs team license

Feature Base runs without a license key Recommended for teams Team license license key on your installation
Board, projects and agents
Docker sandbox for each run
Audit log view
Workspaces One Multiple
SSO providers One Unlimited
Users with roles —
Audit log export (CSV, JSON, SIEM) —
Helm chart for Kubernetes —
Priority support —

The team license includes updates while the license is active. Licensed per team; the price depends on team size.

What you pay for

Octopus team license Priced by team size.
Claude usage Billed by Anthropic to your own key.

We size the license with you in the demo.

Get a license quote

Book a demo

See Octopus run on your own backlog

30-minute demo. Bring a real ticket; we run it through the pipeline live.

  • 30 min
  • Live run
  • Your ticket
  • 01 Your ticket goes from analysis to pull request on screen
  • 02 QA evidence: annotated video, screenshots and report on the task
  • 03 Installed on your servers: we map the setup to your infra and size the team license

Prefer email? [email protected]

FAQ

Questions before the call

Which AI model does it use?

Claude. You bring your own Anthropic key, so usage bills to your account and your limits apply.

Where does my code run?

Every run gets its own isolated sandbox with a short-lived API key scoped to that run. Everything runs on the servers where Octopus is installed: yours.

Can we customize the pipeline?

Yes, per board: stages, the agent on each stage, its prompts, the tools it can use (with read or write scopes) and the automations that move work forward.

How does QA get evidence?

The testing agent attaches annotated videos, screenshots and a test report to the task. QA reviews the evidence instead of reproducing every case by hand.

How long to get started?

One command, octopus init, brings up the whole stack with Docker Compose on your server. For Kubernetes there is a Helm chart.

How is it licensed for teams?

Per team, with a license key installed on your own servers. It turns on multiple workspaces, SSO providers, audit log export, the Helm chart and priority support. Model usage goes on your own Claude key.

Octopus demo

Book a 30-minute demo

Leave your email and we'll write to you from [email protected] within one business day.

Prefer email? Write to [email protected]