Delivery approach · Governed agent workflows

From coding assistants to a governed engineering workflow.

I help engineering teams define how AI agents receive context, carry out bounded tasks, and produce changes that can be independently checked. Start with a pilot on representative engineering work, with clear ownership and acceptance criteria.

A useful agent workflow starts before the first prompt and ends after the code is written. The team needs to know what the agent may change, which evidence will establish success, and who takes over when the work cannot proceed.

This is a proposed delivery approach. The example below is illustrative, not a client case study or a claim of measured production results.

From an agreed brief to a reviewed change
  1. 01Human-approved brief

    Agree the task and acceptance criteria.

  2. 02Context and permissions

    Provide relevant instructions and scoped access.

  3. 03Bounded implementation

    Work within the agreed files and tools.

  4. 04Independent checks

    Test behaviour, security and architecture.

  5. 05Human review

    Review the change and its evidence.

  6. 06Protected merge

    Require passing checks and authorised approval.

If checks fail: a bounded retry, or escalation to the responsible engineer. Unresolved failures block progress to merge.

Give agents the right context

Define the intended behaviour, constraints, relevant architecture and acceptance criteria before delegating. Keep repository instructions and reusable skills versioned. Use connected tools, including MCP where appropriate, to expose only the context and actions the task requires.

The team owns the brief. The agent can flag ambiguity and propose changes, but cannot silently expand scope or weaken the criteria it will be judged against.

Choose the right execution path

Work alongside an engineer

Use GitHub Copilot interactively when the task benefits from discussion, changing context or close supervision. The engineer guides the work and reviews the proposed changes.

Delegate a bounded task

For suitable repeatable work, configure GitHub Agentic Workflows to run coding agents in GitHub Actions, with scoped permissions and controlled outputs. The workflow prepares changes for review within the agreed boundaries.

Both paths need ownership, access controls and merge rules. Separate planning, implementation and checking where that improves the result; choose roles around the task rather than requiring a fixed number of agents.

Tooling is selected for the pilot and the team's environment. See GitHub's Agentic Workflows documentation for the current execution model and controls.

Check outcomes independently

Define checks before implementation: functional and integration tests, relevant security analysis, architecture review and representative evaluation tasks. An agent's own statement that it has finished is not acceptance evidence.

Protect the agreed acceptance checks from being weakened to obtain a pass. Keep the resulting diff, test outcomes and execution trace available for review. Use repeatable evaluation runs to understand which tasks the workflow handles reliably and where human intervention is still needed.

Handle failure deliberately

Set a retry and execution budget. Record progress so a failed run can be inspected and resumed safely. Limit access to tools and secrets, isolate execution, and make repeated actions safe where possible.

If an agent reaches a permission boundary, cannot resolve a failed check or discovers a conflict in the brief, it stops and escalates with the relevant evidence. A responsible engineer owns exceptions, merge and release decisions.

Illustrative workflow

Add input validation to a fictional service

This scenario shows the process, not a completed demonstrator or client result.

  1. Agree behaviour. The brief defines accepted and rejected inputs, the expected error response and existing behaviour that must remain intact.
  2. Bound the change. Permit changes to the relevant handler and implementation tests. Keep secrets, deployment and unrelated files outside the task.
  3. Implement and check. The agent proposes a patch. Independent acceptance tests check valid, invalid and boundary inputs alongside the existing regression suite.
  4. Review the evidence. The PR contains the diff, test results and unresolved questions. Failed checks trigger a limited retry or escalation; the agent cannot waive them.
  5. Decide on merge. An authorised reviewer checks the result against the brief. Required checks and repository rules govern whether it can be merged.

Start with a measurable pilot

Choose one repository and one workflow with representative tasks and an accountable owner. Agree a baseline, access boundaries and acceptance criteria, configure the workflow, then run and inspect the results before deciding what to extend.

What we put in place

A versioned brief and context, scoped permissions, an execution workflow, independent checks, failure handling and a reviewable record of each run.

What we measure

Accepted-task success, review and rework effort, human interventions, regressions, and time and cost per accepted result. Compare these with the agreed baseline.

The output is a working pilot and an evidence-based recommendation to extend, revise or stop. Scope and commercial terms are agreed for your environment; production readiness and savings must be established through the work.

Discuss your project Read the assessment outline