Boards with your own stages
Name the columns after how your team ships. Each stage decides which agent runs and who signs off.
Octopus · by Mister.D
Every task moves through analysis, development, testing and review with a Claude agent on each stage, an auditor that can block the work, and a person on your team who signs off before anything merges.
Checkout fails on Safari
Workflow automation
Define the stages your team already uses. Octopus puts a specialized agent on each one, an auditor on every handoff and a human at the gates that matter. A ticket goes in; a reviewed, tested, deployed change comes out.
From the board, the API or an automation. It enters the first stage of a pipeline you configured, with its own agents and rules.
It checks the code the ticket touches, writes acceptance criteria and turns them into concrete test cases before anyone writes a line.
Files to change, subtasks, risks. The auditor agent approves it, asks for changes or blocks it with a reason written on the task.
Each run gets its own container or microVM and a short-lived API key scoped to that run. The agent opens a branch, commits and opens the pull request.
The testing agent executes every case, captures screenshots and records a video annotated with each step and result. All of it lands on the task.
The auditor reviews the PR, a person approves at the review gate and the merge triggers the deploy through GitHub Actions.
Automations
Triggers watch the board: new task, stage change, comment, PR merged or a schedule. Conditions narrow them. Actions assign an agent, move a stage, notify or run a pipeline.
QA on entry
Ship on merge
Weekly report
Name them, order them, add or remove them. Each board runs its own process.
Each stage has its own prompt, tools and MCP permissions, down to read or write per tool.
Pick the stages where a person signs off. Nothing merges past a gate without it.
Events or cron start the work. Nobody has to remember to push the ticket forward.
Testing & QA
Octopus reads the ticket and its acceptance criteria, writes the test cases, runs each one in a real browser and attaches the recording to the task. Each step is marked on the video: what was clicked, what was typed, what was asserted, and where it broke. QA watches, then approves or comments.
The testing agent reads the ticket and the acceptance criteria, then writes the cases, edge cases included, each with a priority.
Every case runs in a real browser inside the run’s isolated sandbox. Clicks, typing, navigation, console and network are all recorded.
An annotated video per case and a pass/fail report land on the task. Failures open a comment with the exact step that broke.
4.5× faster per ticket
Features
Sixteen pieces that take a task from the board to production, with an agent on every stage and a person on every gate.
Name the columns after how your team ships. Each stage decides which agent runs and who signs off.
Analysis, planning, development, testing and review each get their own prompt, tools and permissions.
A second agent reads the output and answers approve, request changes or block.
Nothing reaches review or production until a person on the team clicks approve.
Agents open a branch, commit, open the pull request and trigger your GitHub Actions deploy through the Octopus GitHub App.
Agents navigate your app, take screenshots and record video with annotations drawn on the exact element that broke.
Every run gets its own container or microVM on Kubernetes, then it is torn down.
GitHub, browser, code sandbox and the Octopus API over MCP. Grant read or write tool by tool.
New task, stage change, comment, merged PR or a schedule fires: assign an agent, move a stage, notify, run the pipeline.
Ask why, give context or redirect it without leaving the ticket.
Screenshots, PDFs, annotated videos and logs land on the task that produced them.
Define the numbers you care about, such as lead time, rework or approvals per stage, and chart them per workspace.
Agents run on Claude with your Anthropic key. Your usage, your bill, your limits.
API keys belong to one workspace with explicit permissions. Each run gets a short-lived key of its own.
SAML single sign-on and passkeys through MisterD ID. Every sensitive action is written to the audit log.
Installs on your infrastructure with Docker Compose or Helm. Team features through a license key. Export and import a whole workspace.
Governance & security
Every run starts in a fresh sandbox, gets a key scoped to one task, touches only the tools you allowed, and passes an auditor and a human before anything ships. Every step lands in the audit log.
Each run gets its own container or microVM on Kubernetes. Nothing carries over to the next one.
A fresh API key per run, limited to its task and workspace, expiring on its own.
Grant GitHub write, browser, sandbox or Octopus API read tool by tool. Read and write scopes are separate.
An auditor agent approves, requests changes or blocks each stage. People sign off at review gates.
Keys minted, tools called, verdicts and approvals, with timestamps and the run that did it.
Sign in through MisterD ID. Workspaces with roles and workspace-scoped API keys.
Octopus installs where you decide: your cloud account, your data center or a single server. Code, sandboxes, evidence and logs stay there. The only outbound call goes to Claude, with your own key.
$ octopus init✓ pulled images (daemon, web, postgres)✓ workspace created, admin key minted→ Octopus ready at http://localhost:8585
A license key turns on the team features in your own installation. Pricing follows team size; we size it with you in the demo.
Licensing
You install Octopus on your own servers, add a team license key to turn on the team features, and connect your own Claude API key. Code, sandboxes, evidence and logs stay with you. The only outbound call goes to Claude.
Your cloud account, your data center or a single server. One command with Docker Compose, or a Helm chart on Kubernetes. Code, sandboxes, evidence and logs stay there.
Install the license key on your own installation and the team features turn on. The key has an expiry date; to renew, you install a new key. No reinstall.
Octopus calls Claude with your own API key. Anthropic bills model usage straight to your account. We add no markup.
| Feature | Base runs without a license key | Recommended for teams Team license license key on your installation |
|---|---|---|
| Board, projects and agents | ||
| Docker sandbox for each run | ||
| Audit log view | ||
| Workspaces | One | Multiple |
| SSO providers | One | Unlimited |
| Users with roles | — | |
| Audit log export (CSV, JSON, SIEM) | — | |
| Helm chart for Kubernetes | — | |
| Priority support | — |
The team license includes updates while the license is active. Licensed per team; the price depends on team size.
We size the license with you in the demo.
Get a license quoteBook a demo
30-minute demo. Bring a real ticket; we run it through the pipeline live.
We will reply within one business day with times. Pick a ticket in the meantime.
FAQ
Claude. You bring your own Anthropic key, so usage bills to your account and your limits apply.
Every run gets its own isolated sandbox with a short-lived API key scoped to that run. Everything runs on the servers where Octopus is installed: yours.
Yes, per board: stages, the agent on each stage, its prompts, the tools it can use (with read or write scopes) and the automations that move work forward.
The testing agent attaches annotated videos, screenshots and a test report to the task. QA reviews the evidence instead of reproducing every case by hand.
One command, octopus init, brings up the whole stack with Docker Compose on your server. For Kubernetes there is a Helm chart.
Per team, with a license key installed on your own servers. It turns on multiple workspaces, SSO providers, audit log export, the Helm chart and priority support. Model usage goes on your own Claude key.