RiskMandate v1.34.15
The business case · Agent harnesses and runtimes

Cloudgeni and Opengeni, by the risk they change

Cloudgeni runs AI agents against cloud infrastructure and delivers every change as a pull request; its documentation says agents do not deploy, cloud access is read-first, and runs are scoped. It runs on Opengeni, an open-source harness whose README describes durable sessions, tool approvals that pause a run for a person, a replayable event log of every step, and execution in a managed sandbox or on a machine you own. For RiskMandate the case is the harness: put an agent with a behaviour policy inside it and the rows that today rest on an instruction or a setting gain a boundary the agent cannot reach. The published Gmail ABP is the worked example.

Disclosure: RiskMandate's lead spoke with one of Cloudgeni's founders on 25 September 2026, and the two companies are setting up a pilot: running the agents behind RiskMandate's Gmail and Calendar behaviour policies inside Opengeni. This case was written before the pilot, from Cloudgeni's and Opengeni's public pages only, read on that date. Nothing from the conversation is on this page; when the pilot produces a measured result, the case changes with the date.

Open source: Apache-2.0 for Opengeni, the harness it runs on; Cloudgeni itself is a hosted service · Cloudgeni (Oslo) · get involved

The deployment: The model's typical deployment, moved to what an infrastructure agent usually looks like without a harness: it changes production on its own, what it did cannot be reconstructed, stopping it is possible eventually, nobody knows how long that takes, and nobody can say what it could reach if it misbehaved.

The model: the RiskGraph Explorer's 49 facts, 49 risks and 10 roles, copied into this site with its provenance; the register below is computed, not written. How.

01 · What it does

In its own words, and nothing more.

  • Runs agents against cloud infrastructure and delivers the result as a pull request, not a deployment. “Agents do not deploy. No direct apply, deploy, or production cloud mutation. Cloud access is read-first. Inventory, findings, resource context, and validation inputs. Changes land as PRs. Your team reviews, merges, and deploys. Runs are scoped. Selected repositories, integrations, tools, and organization data only.” · docs.cloudgeni.ai, read 25 September 2026
  • Keeps deployment credentials out of the agent's hands by design. “Deployment-grade credentials should stay in customer CI/CD. Cloudgeni can generate and validate the proposed IaC; the final plan/apply or deploy step remains customer-controlled.” · docs.cloudgeni.ai, read 25 September 2026
  • Records what a run did: messages, tool output, credential use and audit events, exportable through the API and the CLI. “Session messages, tool output, scan state, credential usage, audit events, and telemetry are persisted for review.” · docs.cloudgeni.ai, read 25 September 2026
  • Runs on Opengeni, an open-source harness that holds the session, the approvals and the log around any agent. “Opengeni is a production-ready agentic service: it runs AI agents that do real work, keeps a session going for hours or days, records every step in a replayable event log, stops for a human when an action needs approval, and puts each session either in a managed sandbox or directly on a machine you own.” · github.com, read 25 September 2026
  • Pauses a run for a person before a risky tool call, and resumes the exact call after the answer. “Tool approvals gate risky actions. Agents can ask structured questions and resume the exact tool call after the answer, even across restarts.” · github.com, read 25 September 2026
  • Stores organisation-specific guidance as versioned policies the agent is given as context: an expectation, in this site's vocabulary, and named as guidance on its own page. “Custom policies are reusable guidance that Cloudgeni applies in analysis and agent workflows.” · docs.cloudgeni.ai, read 25 September 2026
02 · The answers it changes

Same deployment, different answers.

The model asks sixteen questions about an agent deployment. A product’s effect is written as the answers it changes, and each change says what kind of change it is: a statement of what is true, an expectation the agent is asked to meet, a setting, or a boundary enforced by something the agent’s grant does not include.

The questionWithout itWith itHow, and who holds it
Can it change things in production, or only read and report?Changes on its ownChanges, with a person approving each onea boundary, held by the customer's reviewers, and their CI/CD. A boundary because the deployment credential is not in the agent's grant: the agent writes a branch and a pull request, and a person merges. It holds for changes that go through the Git path. A Connected Machine, which Opengeni drives directly, reaches whatever that machine reaches, and is the customer's own choice. “The product is not designed around the agent writing directly to production infrastructure.” · docs.cloudgeni.ai, read 25 September 2026
Could you reconstruct what it did last Tuesday?NoFullya setting, held by the operator of the harness. For what the session did: every event in the log, replayable. A setting rather than a boundary because the log is the operator's to keep, export or delete, and because what a Connected Machine did outside the session's tool calls is not in it. “Every event lands in Postgres. Live streams backfill from it, so a browser reload, a new client, or an audit replays the same history.” · github.com, read 25 September 2026
If you had to stop it right now — could you?Yes, eventuallyYes, one actiona setting, held by the operator of the harness. One action, from the session: a person interrupts, or the run pauses itself at an approval. A setting because the harness's own operator holds it, and an agent on a Connected Machine is stopped by stopping the machine. “The agent keeps working until it completes the goal with evidence, pauses with a rationale, or a human interrupts.” · github.com, read 25 September 2026
How long would stopping actually take?Don't knowMinutesa setting, held by the operator of the harness. Minutes, for a person watching the session. The pages do not state a time; the interrupt is a control in the session's own interface. “Give agents work from the web app and follow along, or call the same session API from your own product and let Opengeni hold the state, history, approvals, and outputs.” · github.com, read 25 September 2026
If it misbehaved at full speed, what could it reach?Don't knowInternal onlya boundary, held by whoever configures the run's scope and sandbox. For a run in a managed sandbox with selected context and task-specific tools, what it could reach is the list it was given, read-first on the cloud side. Whether that list is internal only is the customer's to say; what changes is that the reach is stated at dispatch rather than discovered afterwards. On a Connected Machine the reach is the machine's. “The agent receives selected context, task-specific tools, and an ephemeral run context.” · docs.cloudgeni.ai, read 25 September 2026
03 · The register, before and after

9 retired, 17 unchanged.

Computed from the model for the deployment above: every risk that holds without it, and every risk that holds with it.

9retired
0new
26 → 17entries on the register

Retired

RISK-8
The operator cannot state what the agent is entitled to reachoperational · SRE / platform on-call, Platform owner
RISK-9
Stopping the agent takes more than one actionoperational · SRE / platform on-call
RISK-13
What the agent did cannot be reconstructedoperational · SRE / platform on-call, CISO
RISK-15
The organisation acts through a system that acts without a person approving each actionoperational · CTO, Board
RISK-17
The scope of the agent's access cannot be reviewedoperational · CISO
RISK-19
The time a stop takes is unknown to the people who would perform itoperational · SRE / platform on-call, Platform owner
RISK-29
The stop capability cannot be compared to any exposure windowoperational · CISO
RISK-33
The organisation could not account for what its agent didoperational · Board
CORP-5
Accountability failure — the organisation cannot demonstrate who decided whatcorporate · CEO, Board
04 · From the operator to the board

Who carries less, and who carries the same.

Each risk is assigned to the roles it belongs to, and each role reports to another until the board. The count beside each role is the entries it holds without the product and with it.

Operators

who run it, and are paged when it goes wrong
SRE / platform on-call10 → 6 −4retired RISK-8, RISK-9, RISK-13, RISK-19
Customer-facing service owner1 → 1

Owners

who own what it may touch, and can say what happened
Platform owner4 → 2 −2retired RISK-8, RISK-19
CISO9 → 6 −3retired RISK-13, RISK-17, RISK-29
DPO6 → 6

Executives

who answer for speed, product, cost and the whole
CTO3 → 2 −1retired RISK-15
Chief Product Officer3 → 3
CFO0 → 0
CEO4 → 3 −1retired CORP-5

The board

who must be able to defend how it is run
Board6 → 3 −3retired RISK-15, RISK-33, CORP-5

At the board: the corporate register

Corporate risks have no facts of their own. They hold while any risk that leads into them holds, so a single product rarely retires one. What it changes is how many reasons the board is being given.

CORP-1
Regulatory exposure — the organisation may be operating outside obligations it is subject tostill holds · risks leading into it: 3 before, 1 after
CORP-2
Customer harm — customers may be affected by decisions or changes nobody madestill holds · risks leading into it: 2 before, 2 after
CORP-3
Operational resilience — the organisation may be unable to restore a service it depends onstill holds · risks leading into it: 1 before, 1 after
CORP-5
Accountability failure — the organisation cannot demonstrate who decided whatretired · risks leading into it: 1 before, 0 after
CORP-6
Loss of control — the organisation operates a system it may not be able to stop or explainstill holds · risks leading into it: 2 before, 1 after
06 · What it adds

Every product is also a new thing in the estate.

A control plane holding the run's material: its own page lists infrastructure metadata, selected repository content, findings, prompts, generated changes, and audit events, and the privacy policy on cloudgeni.ai says scan data is typically retained for 2 years. A cloud read credential per account: on AWS its setup page creates a user with the managed ReadOnlyAccess policy; on Azure the manual setup assigns Reader, Cost Management Reader, Security Reader, Log Analytics Reader and Storage Account Key Operator Service Role, the last of which is not a read role by name and is not explained on the page read. The harness's own components, which the README lists for a self-host: Postgres, NATS, Temporal, object storage, the API, workers and the web app. And, in local development, the README's default that the agent runs commands directly on the developer's machine rather than in a container, with the sandbox as an option to turn on.

Free, but not free

What adopting it takes, before any of this is true.

An open-source project costs nothing to download and something to adopt. Every change above depends on the work below, and most of it is customisation to your own deployment.

  • Decide the sandbox: a managed sandbox for the run, or a Connected Machine, which is driven directly and reaches what the machine reaches. The boundaries in this case assume the sandbox.
  • Write the approval rules: which tool calls pause for a person, and who that person is. An approval is a person clicking, and the residual risk of that is the CISO's to accept.
  • Keep deployment credentials in the customer's CI/CD, as the trust page says, and check that no run context carries one.
  • Decide who reads the event log, how long it is kept, and what is redacted before it is exported; the log is the evidence and also a copy of what the agent saw.
  • Grant the cloud read roles the setup page names and no more; on Azure, ask what the Storage Account Key Operator role is for before assigning it.
  • If self-hosting, run the reference deployment and own its components; if using the managed service, read what the control plane holds and where.
Where its pages disagree

Published unresolved, for Cloudgeni (Oslo) to settle.

Each pair was read on the same day. We have not tested which is true, because that would mean testing somebody else’s system.

  • does a change ever merge on its own. The marketing site, in its own text, says Autonomy is a policy — routine changes merge on their own, critical paths wait for review. The documentation's trust page says Changes land as PRs. Your team reviews, merges, and deploys, and the AI DevOps page says the product is not designed around the agent writing directly to production infrastructure. Whether a routine change can merge without a person is left to Cloudgeni to state. cloudgeni.ai · security-commitment
  • which clouds. cloudgeni.ai/llms.txt, last updated 15 September 2025 by its own header, lists AWS and Azure. The documentation's FAQ lists AWS, Azure, GCP, Kubernetes, OCI, Exoscale and OpenShift. Both were served on 25 September 2026. llms.txt · faq
What it gives an Agent Behaviour Policy

Four objects, and which of them a run fills.

The pilot's question, stated in the ABP's terms. RiskMandate's published behaviour policy for Claude's Gmail connector has six rows, a mandate of read and draft, never send, and four rows of unbounded excess: four things the agent can do that nobody asked for, with nothing out of its reach in the way. A harness does not change what the connector grants. What it changes is the barrier column. The table takes each Gmail row as it stands today, from the vault, and says what Opengeni's pages would let a run inside it do to that row. Every entry in the third column is a projection from the pages, not a measurement; the pilot is the measurement.

The ABP objectWhat the product suppliesHow it becomes the ABPWhat stays open
send.message.world · send mail as the person, to anybodyToday: a setting. Claude's per-action approval prompt, which the person's own account can switch to always allow. In excess, refused by the mandate, unbounded.Inside the harness the send becomes a tool call that stops for a human when an action needs approval, held by the operator of the harness, outside the agent's account. That is a boundary. The row leaves unbounded excess.The approval is a person clicking; the residual is the CISO's. And the harness must hold the Gmail credential itself, so the send goes through it and not around it.
create.schedule.tenant · create something that outlives the sessionToday: a setting, behind the same approval prompt. In excess, refused, unbounded.A run inside Opengeni has its own sessions, goals and schedules, owned by the harness and its operator. Whether the agent can still create a schedule in Claude's own tenant depends on which surface the pilot connects the mailbox through; if the agent's only surface is the harness, the row leaves the grant.Open until the pilot decides the surface. The case counts it as bounded only if the Claude surface is not in the run.
read.credential.host · read credentials on the hostToday: nothing in the way, inferred. In excess, refused, unbounded.In a managed sandbox there is no host of the person's to read: an ephemeral run context, with task-specific tools. The row leaves the grant for a sandboxed run.On a Connected Machine the row is the machine's, and the README says the machine is driven directly. The pilot should run in the sandbox.
read.record.history · read shell history and past sessionsToday: nothing in the way, measured. In excess, unstated, unbounded.The same: the sandbox holds no history of the person's. What it holds is the harness's own replayable log, which is a different record with a different owner.The log is a copy of what the agent read, including mail. Who may read the log becomes Legal's row, as it was for the proxy in the role-ownership article.
authenticate-as.credential.tenant · act in the mailbox as the personToday: a boundary, the OAuth consent, revocable by the person. In excess, unstated, bounded.Unchanged in kind. The credential moves from the person's Claude session to the harness, so revoking it means revoking the harness's grant, and the person should know that.Where the harness keeps it, and who can read the harness's secrets, is the trust page's secrets row in Opengeni's governance layer; the pilot reads that page.
read.message.tenant · read the mailboxToday: a boundary, the consent's read line. Aligned with the mandate.Unchanged. The mandate wants it; the harness does not narrow a read scope that Google issues for the whole mailbox. The instructions in the run say which mail is for the job; that stays an expectation.The same fact this site records on every mail connector: the scope covers the account, not the business process.
The Calendar policies, in draftTheir dangerous row is an edit that Google documents no way to restore.A harness that pauses before the edit tool call, and whose log records the event as it stood before the call, is the barrier the Calendar article asked for: a before-image of every event the agent edits, kept where the agent cannot reach. Opengeni's replayable log is where that before-image would live.Only if the tool reads the event before it writes it, and the log keeps the read. That is a design decision for the pilot, and it is the one worth making first.

The projection, if the pilot runs in the sandbox and the mailbox is reached only through the harness: the Gmail policy's unbounded excess goes from four to nought, its excess stays at five, and two residual risks appear where none was recorded, the approval that is a person clicking and the log that is a copy of mail. The register above is the model's; these counts are the vault's, and the pilot will replace them with measured ones.

From the buyer’s desk

What it does, for the person who signs.

Cloudgeni's pages are written for the platform engineer. The person who has to sign for an agent that touches production wants three things from a harness: that it cannot deploy, that it stops for a person, and that afterwards somebody can say what it did. This is the run from that desk, in the order the trust page gives it.

One run, from the task to the pull request 01 A task starts A prompt, aschedule, a webhook,a scan result or adrift item createswork. 02 The scope isresolved Organisation,workspace,repository,integration andactor, beforeanything isdispatched. 03 A worker runsthe agent In a managed sandboxor on a machine youown, with selectedcontext,task-specific toolsand an ephemeral runcontext. 04 It stops for aperson A risky tool callpauses the run; theanswer resumes theexact call, evenacross restarts. 05 Every step isrecorded Messages, tooloutput, credentialuse and auditevents, in areplayable log. 06 The result isa pull request Your team reviews,merges and deploys.The deploymentcredential neverleft your CI/CD. WHO READS IT The platform engineer Findings turned into reviewed IaC, drift intopull requests, and an agent that works withrepo and cloud context. The CISO Read-first cloud access, a scoped run, anapproval gate, and a log that can be exportedto the SIEM. The person who signs An agent whose excess is bounded by the harnessrather than by its instructions: the ABP'sbarrier column moved, with evidence per run.
Drawn from the product’s own pages, read on the date at the top of this case. Nothing was run.

It cannot deploy, because it does not hold the credential.

This is the strongest line on the trust page and the reason the case gives it a boundary rather than a setting. A control that works because the agent's grant does not include the thing is the only kind this site counts as a control.

The harness is separate from the agent, and open source.

Opengeni's README says it is not the agent; it is everything the agent needs around it. That is exactly the shape a behaviour policy wants: the agent's grant on one side, the barriers on the other, held by somebody else.

An approval is a boundary and a residual risk at once.

The run pauses for a person. The person can be worn down, and the pause is only as good as the rule that decides which calls pause. The case counts the boundary and names the residual.

The log is the evidence and a copy of what the agent saw.

Replayable history is what makes last Tuesday reconstructable. It is also where a customer's mail or repository content now lives, for as long as the operator keeps it. Both are true; the second belongs to Legal.

Two ways to run, two different pictures.

In a managed sandbox the reach is what the run was given. On a Connected Machine the reach is the machine's. Every boundary in this case is for the first; a buyer should ask which one they are getting.

The pilot, and the case as a product

What the two companies can make together

RiskMandate's offer to Cloudgeni, and to every product on this section, is the same: the business case in the register's terms, computed and dated, with the behaviour policy's counts moving as the proof. These are the notes for this one.

The first deliverable is a measured Gmail ABP inside Opengeni.

Take the published Gmail policy, run its agent in the harness, and recompute the delta from what the harness actually bounded. The projection above says unbounded excess goes from four to nought; the measurement will say what it really does, row by row, with the evidence in the run's own log. That document is the pilot's output and the case's next version.

The Calendar before-image is the demonstration to lead with.

Google keeps no version of an edited event. A harness that pauses before the edit and keeps the event as it was in its replayable log is a control nobody else on this site has, for a risk that stops calendar agents being adopted at all. It is small, it is specific, and a buyer can see it in one screen.

Sell the harness by the answers it changes.

Five answers move on this page, each with its kind and its holder. A page like this for each agent shape Cloudgeni's customers run, computed from the same model, is a sales document a CISO can forward to the person who signs. RiskMandate writes them; Cloudgeni corrects them; both carry a date.

A vault per run, or per agent, as the evidence pack.

Opengeni already keeps a replayable log. An encrypted vault the customer holds the keys to, with the behaviour policy beside the run's log and the delta recomputed per version, is the artefact an auditor or an insurer is handed: a read key, not a login. This site's policies are built that way; the pilot can package its first result the same way.

Open source the way open-source.sgit.ai describes it.

Opengeni is Apache-2.0 with a managed service on top, which is packaging so long as the self-host and the service run the same code. The test from open-source.sgit.ai is one question, asked at each release: does the managed service contain anything the repository does not? Say the answer on the download page, and the trust page gets stronger.

What not to claim.

That the harness makes an agent safe. It bounds what the grant gives; it does not narrow the grant. The mail scope still covers the mailbox, the read is still whole, and the instructions are still an expectation. The case is strong because it says which rows move and which do not.

The position these rest on is published at open-source.sgit.ai, with its counter-cases; the vault pattern is the one this site’s own behaviour policies use, described on how it works.

What this does not claim

The limits of the case, stated by the case.

  • That it was tested. Nothing was installed or run; every change rests on the pages quoted beside it. The pilot is where a measured result will come from, and it has not started.
  • That a Connected Machine is bounded. The README says machines are driven directly; every boundary above is for a run in a managed sandbox, and the case says so on each line.
  • That custom policies are a control. Their page calls them guidance; in this site's vocabulary that is an expectation, and it is listed with what the product does rather than with what it changes.
  • Anything from the conversation with the founder. The disclosure says it took place and that a pilot is planned; the page uses the public pages only.
  • That the register is complete. It is one model, for one stated deployment. A different deployment changes the answers, and so the case.

Next. Published, and sent to Cloudgeni the same day. The pilot's first deliverable is the Gmail behaviour policy measured inside Opengeni, with the counts below recomputed from what the harness actually bounded; that result replaces the projection in the ABP section, with its date.

The same method, for your product

Make the case in the register’s own terms.

If you build a security product for agents, the case for it can be written the same way: what it does in your own words, the answers it changes, and the register before and after. If a case here is wrong about you, tell us and it changes with a date.