Five decisions from release freeze to validated recovery

Source-control platform outage during an active release

A source-control outage becomes more than an availability problem when a release is already moving. A tag may or may not have been created, automated checks can continue after repository access disappears, local clones can disagree, and an artifact may reach staging before anyone names the person who can stop the process. This tabletop asks a team to control that uncertainty without treating a provider outage as proof of a security compromise.

Use this guide beside the Interactive Source-Control Platform Outage Rehearsal. The interactive workspace supplies five injects, three decision paths per inject, a timer, readiness meters, and an after-action export. This page supplies the facilitator objective, room roles, evidence structure, neutral prompts, and recovery criteria needed to make those choices meaningful. Keep the exercise fictional or sanitized; do not enter real credentials, tokens, repository names, customer details, or active incident information.

Set one decision-centered objective

By the end of the session, the team should be able to name who can freeze and resume an active release, reconstruct the last trusted release state from more than one evidence source, choose a bounded continuity path, assess credential exposure without assuming compromise, communicate a cause-neutral status, and restore service through validation gates with durable follow-up owners. The exercise is not a test of a particular source-control product, CI system, command, or branching model.

Define success in observable terms. Examples include invoking a release hold within five exercise minutes, recording one authoritative release owner, distinguishing verified evidence from a developer's local recollection, and refusing to resume production deployment until artifact provenance and rollback behavior have been checked. If the room debates root cause for most of the hour, return to the decision that must be made with the facts currently available.

Assign the roles and authorities before the first inject

  • Facilitator: reads the injects, protects the clock, asks neutral questions, and keeps the room focused on decisions rather than product trivia.
  • Exercise incident lead: maintains the operating picture, sequences response work, and identifies where another role owns approval.
  • Release authority: can freeze, resume, roll back, or cancel the release under the organization's current process. If this authority is unclear, record that as an exercise finding.
  • Engineering and CI owner: explains repository, runner, build, artifact, deployment, and rollback evidence without treating one system as the entire record.
  • Identity or security owner: evaluates token and credential exposure, preserves access evidence, and recommends proportionate restrictions or rotations.
  • Business and communications owner: coordinates leadership, support, product, partner, and vendor updates using confirmed facts and a stated next-update time.
  • Scribe or observer: records facts, assumptions, decisions, owners, tradeoffs, deadlines, and the evidence that could change each decision.

One person can cover several roles in a small team, but do not leave authority implicit. Ask who can stop a production change after hours, who accepts delay against a customer commitment, and who can authorize an emergency repository, runner, or credential path. The answer should be a role and current authority, not merely the most senior person in the room.

Run the five-stage agenda

  1. 0-10 minutes - Detect and establish release control. The primary source-control platform fails while an approved release is moving through automated checks. Some jobs still run and an artifact has reached staging. Ask whether the release is frozen, what the freeze covers, who invoked it, which actions may continue for evidence preservation, and when the decision will be reviewed.
  2. 10-20 minutes - Triage repository and CI evidence. Provide partial CI logs, runner records, an artifact registry, deployment telemetry, and conflicting local clones. Ask the team to identify the intended commit, approvals, tag state, executed jobs, produced artifact, and last trusted deployment. Require every claim to be marked confirmed, inferred, conflicting, or unavailable.
  3. 20-32 minutes - Contain risk while considering continuity. Introduce an approved mirror, cached source on controlled runners, a last-known-good artifact, and active automation tokens. Ask which fallback is allowed, how freshness and provenance will be checked, what credentials must be limited, and what condition ends the exception. Remind participants that service unavailability does not establish credential theft.
  4. 32-42 minutes - Coordinate stakeholder and vendor communications. Product, operations, support, and leadership want a release decision while the provider requests timestamps and request identifiers. Require one status owner, confirmed impact, unresolved questions, stakeholder actions, the provider evidence package, and a next update. Do not let the team call the event a compromise without supporting evidence.
  5. 42-55 minutes - Restore in stages and validate. Repository access and CI triggers return. Ask what must be reconciled before merges resume, which artifact will be tested, how protections and tokens will be checked, what canary or limited deployment is appropriate, and which failure condition triggers rollback. Use the last five minutes for the AAR and improvement owners.

Build a release evidence ledger

The ledger should explain what the team believed at each decision point and why. It is not a request to paste real logs into the exercise. Use generic artifact names and record only fictional or sanitized values. Give each row an evidence owner, collection time, confidence label, and retention location.

  • Release intent: approved change, expected commit, branch, tag, approver, change window, release owner, and rollback target.
  • Repository state: last verified remote state, protected-branch events, tag evidence, mirror freshness, and differences among controlled clones.
  • CI and runner state: job identifiers, trigger source, runner identity, start and stop times, checks completed, logs preserved, and queued work cancelled or allowed.
  • Artifact provenance: build source, checksum, signing or attestation evidence where used, registry time, promotion history, and environment reached.
  • Deployment state: target environment, deployed version, health checks, configuration changes, canary population, and rollback readiness.
  • Access and token state: token owner, intended scope, last observed use, active sessions, emergency access, restrictions applied, and rotation decision.
  • Provider evidence: incident notice, support case, affected functions, timestamps, request identifiers, provider statements, and the next vendor update.

Do not declare the newest timestamp or closest developer clone authoritative by convenience. Confidence comes from reconciling independent records. When evidence conflicts, preserve both versions, identify who will resolve the conflict, and state which production action remains blocked in the meantime.

Ask release-authority questions that expose real gaps

  • Who can freeze merges, builds, artifact promotion, and production deployment, and are those separate authorities?
  • Does a freeze stop evidence collection, or only actions that can change release state?
  • Who may approve an alternate repository, runner, artifact, or manual deployment path?
  • What evidence is required before the release authority can resume the original pipeline?
  • Who accepts the business impact of delay, and who accepts the residual risk of a bounded fallback?
  • What event, deadline, or failed validation automatically returns the release to a frozen state?

A good answer names one accountable role, its source of authority, the exact action approved, and a review trigger. "Engineering will decide" is not enough when engineering, operations, security, and product each control a different part of the release path.

Bound credential risk without assuming compromise

An outage can limit access to audit data and make token status uncertain, but uncertainty is not evidence of theft. Begin with an inventory: which source-control, CI, registry, signing, deployment, webhook, and emergency credentials could participate in the release; where they are valid; and what activity can still be observed. Preserve available authentication, runner, and provider evidence before broad changes erase useful context.

Apply proportionate controls. Pause high-risk automation that no longer has an accountable release path, reduce scope where the platform permits it, require fresh approval for emergency use, and monitor known integration points. Revoke or rotate immediately when evidence, exposure, policy, or inability to contain use supports that action. Otherwise, document why a credential remains active, who owns the risk, what is being monitored, and when the decision expires. After restoration, reconcile actual token use and close every temporary credential or exception. Avoid both extremes: leaving all access untouched because no compromise is proven, or destroying continuity and evidence by revoking everything solely because the provider is unavailable.

Worked decision record

At exercise minute 24, the release authority keeps production deployment frozen but permits a controlled validation build from the approved mirror. The decision is based on a matching reviewed commit identifier, preserved CI approval record, and a mirror freshness check. The team has not confirmed the primary tag state, so the build cannot be promoted. Engineering owns the isolated build; Security owns the automation token inventory and limits the mirror runner to the required repository and registry path. The accepted tradeoff is a delayed customer change window. Revisit the decision when the provider restores tag and audit visibility or in 20 exercise minutes, whichever comes first.

At minute 37, Communications sends a cause-neutral internal update: the source-control provider is experiencing an availability incident; the release remains frozen; no credential compromise or customer impact is confirmed; validation is occurring through an approved isolated path; Product owns partner coordination; and the next update is due in 15 minutes. This record is useful because it preserves facts, uncertainty, authority, tradeoff, owners, and revisit conditions without pretending the team knows the provider's root cause.

Use a staged-restoration checklist

  1. Reconcile: compare commits, branches, tags, approvals, protected-branch events, runner activity, queued jobs, token use, artifacts, and deployments against the evidence ledger.
  2. Re-establish control: confirm branch protections, required reviews, release permissions, webhook destinations, runner trust, token scope, and the named release authority.
  3. Choose the source: identify the verified commit and artifact path. Rebuild when provenance is incomplete; do not promote an artifact merely because it already exists.
  4. Validate outside production: run required checks in an isolated or staging environment, compare checksums and expected behavior, and preserve the resulting evidence.
  5. Canary: use a limited deployment or equivalent controlled exposure with health criteria, monitoring owner, observation period, and an executable rollback.
  6. Resume deliberately: require the release authority to state which gates passed, what residual uncertainty remains, who accepts it, and what would stop the release again.
  7. Retire exceptions: close alternate repository paths, emergency runners, temporary permissions, tokens, manual steps, and communication channels that are no longer needed.
  8. Record disposition: document the final released or cancelled state, affected windows, provider findings, stakeholder notices, and follow-up evidence owners.

Score the AAR with behavior-based evidence

Score each dimension from 0 to 3: 0 - not observed, 1 - improvised, 2 - defined but inconsistent, and 3 - clear and repeatable. Add one sentence of observed behavior beside each score. The result is an exercise aid, not a compliance grade or maturity certification.

  • Release control: did the room name a freeze scope, release authority, allowed evidence actions, and explicit resume criteria?
  • Evidence discipline: did participants reconcile repository, CI, artifact, and deployment records while labeling facts, assumptions, conflicts, and unavailable evidence?
  • Continuity governance: did the fallback have freshness and provenance checks, narrow authority, monitoring, stop conditions, and a retirement deadline?
  • Credential judgment: did the team investigate exposure and token use without treating outage as compromise or using uncertainty as a reason for inaction?
  • Communication quality: did updates separate availability, release integrity, credential risk, and customer impact while naming an owner and next-update time?
  • Restoration quality: did recovery reconcile state, restore controls, validate outside production, use a limited gate, preserve rollback, and close exceptions?
  • Improvement ownership: did each priority action receive one accountable owner, target date, expected completion evidence, and scheduled review?

Choose no more than three durable improvements. Strong candidates include documenting after-hours release-stop authority, testing a read-only repository or artifact evidence path, defining mirror freshness criteria, retaining CI evidence outside the primary platform, reviewing automation-token ownership, and rehearsing a controlled restoration. An action is not complete when a document is drafted; it is complete when the named evidence proves the new process can be used.

Return to the facilitator guides hub for shorter formats and general room-running guidance. When the objective and roles are ready, open the interactive rehearsal and use its timer, choices, facilitator notes, and AAR export to run the decision path.

Run the source-control outage rehearsal