PluginBench
Skill
Pass
Audit score 90

grok-delegate

amelnagdy/delegate-skills

Delegate coding tasks to Grok Build CLI as a background implementer, then review and land the diff yourself.

What is grok-delegate?

This skill lets you hand a bounded coding task to the Grok Build CLI (`grok`) as a separate implementer while you stay the orchestrator and reviewer. You write the brief, Grok does the implementation under an explicit autonomy profile, and you verify the diff and commit it yourself. Use it when the user wants to delegate work to Grok (phrasings like "have Grok do X" or "delegate this to Grok") and will review and land the result.

  • Dispatch a coding task brief to Grok Build CLI and capture the run in a structured result.json
  • Run Grok under a scoped autonomy profile (workspace-write by default) so changes stay isolated
  • Block until Grok finishes, then return control to you for review and verification
  • Support resuming a previous Grok session with delta briefs to iterate on feedback
  • Provide touchedFiles and full implementation reports so you can audit the diff before committing

How to install grok-delegate

npx skills add https://github.com/amelnagdy/delegate-skills --skill grok-delegate
Prerequisites
  • Grok CLI installed globally (`npm i -g @xai-official/grok`) and authenticated (`grok login` or `XAI_API_KEY` set)
  • Grok Build beta access (eligible xAI subscription)
  • Node 18 or later
  • Git repository with write access
  • Orchestrating agent capable of running shell commands and reading files
Claude Code
Cursor
Windsurf
Cline

How to use grok-delegate

  1. 1.Write a clear, self-contained brief describing the task, current state, what to change, project gate commands, and report contract (see references/writing-the-brief.md)
  2. 2.Run the relay helper: `node "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo` and wait for completion
  3. 3.Read the result.json file and review Grok's full report in the finalMessage field
  4. 4.Re-run your project's gates (test, lint, build commands) yourself to verify they pass
  5. 5.Read the diff against the brief to confirm Grok did exactly what was asked, no scope creep, no omissions
  6. 6.Commit the verified work yourself with a clear commit message; if changes are needed, send a delta brief with `--resume-last`

Use cases

Good for
  • Delegate a refactoring task to Grok while you review the diff and handle the commit
  • Run a queue of implementation tasks through Grok Build in the background, reviewing each one
  • Hand off a bug fix or feature to Grok with a detailed brief, then verify it passes your project's gates before landing
  • Iterate on Grok's output by sending delta briefs with `--resume-last` without restating the whole task
  • Use Grok as a trusted background implementer in CI/CD or local workflows where you control the final commit
Who it's for
  • Developers using Claude Code or Cursor who want to delegate implementation work to Grok Build
  • Teams running coding tasks through Grok while keeping a human reviewer in the loop
  • Engineers who want to hand off bounded tasks to an AI implementer but retain control over what lands in the repo

grok-delegate FAQ

When should I use this instead of writing code inline?

Use this when the task is substantial enough that delegation overhead is worth it, the user explicitly asks to delegate to Grok, or you want to keep the implementation separate for review. Don't use it for small tasks you can do directly.

What if Grok's output doesn't pass my project's gates?

Re-run the gates yourself (don't trust Grok's self-report), read the diff carefully, and send a delta brief with `--resume-last` describing what failed and what to fix. Iterate until it passes.

Does this skill commit the code for me?

No. You always commit. The skill dispatches the task, captures the result, and gives you the diff to review. You decide whether to commit after verifying the gates and the changes.

What if the grok CLI is not installed or authenticated?

Run `npm i -g @xai-official/grok`, then `grok login` (or set `XAI_API_KEY`). Confirm with `grok version`. The relay will fail with a clear error if grok is unavailable.

Can I resume a Grok session if the first run didn't finish the task?

Yes. Use `--resume-last` with a delta brief (only the new instructions, not the whole task) to continue the previous session without restating what's already done.

Full instructions (SKILL.md)

Source of truth, from amelnagdy/delegate-skills.


name: grok-delegate description: >- Delegate a coding task to the Grok Build CLI as a background implementer, then review its diff and land it yourself. Use this whenever the user wants to hand implementation work to Grok — phrasings like "have Grok do X", "delegate this to Grok", "run it through Grok", "use Grok Build to implement/fix/refactor", or "have grok CLI do this" — or to run a queue of coding tasks through Grok while staying the reviewer. Prefer it when the user will review the diff and commit it themselves. DO NOT USE for tasks small enough to do inline, or when the user wants the code written directly without delegating. license: MIT compatibility: Requires the grok CLI (Grok Build) installed and authenticated (grok login, or XAI_API_KEY; beta access needs an eligible xAI subscription), Node 18+, and git. The orchestrating agent must be able to run shell commands and read files. Shell examples assume bash/zsh (macOS/Linux, or Git Bash/WSL on Windows). metadata: version: 0.5.0

Grok Delegate

For a trusted repository rejected by Git's ownership check, the relay supports --trust-git-root <exact-worktree-root>. This opt-in affects only relay Git checks, without persistent Git config or Grok permission changes. See dispatch and poll.

You are the orchestrator. This skill lets you hand a bounded coding task to a separate implementer — the Grok Build CLI (grok) — then review what it produced and land it yourself. You write the brief and own the judgment; Grok does the typing under an explicit autonomy profile; you verify and commit.

Nothing here is specific to one orchestrating agent. The loop needs only the ability to run a shell command and read a file, so it works the same whether you are Claude Code, Cursor, OpenCode with a selected model, or any comparable agent. (It is designed for Claude Code and Cursor; treat other orchestrators as designed-for, not yet proven.)

When NOT to use this

  • The task is small enough to just do inline — delegation overhead is not worth it.
  • The grok CLI is not installed, not authenticated, or the account lacks Grok Build beta access.
  • You want to write the code yourself, or you only need a review without an implementer run.

Prerequisites (check once)

  1. grok version succeeds. If not, install on any platform with npm i -g @xai-official/grok (or use the installer from xAI's official Grok CLI docs) and authenticate (grok login, or grok login --device-auth on headless hosts, or set XAI_API_KEY).
  2. Confirm which grok is on PATH. command -v grok shows the active binary and grok version its version — the relay records the version it ran into result.json, so a stale binary is visible after the fact.
  3. You are in (or will point --cd at) the target git repository.

The loop

Run these five steps per task. Steps 1, 4, and 5 are your judgment; 2 and 3 are mechanical.

1. Write the brief

Grok sees only the text you send — no orchestrator chat history, no shared context. Everything the task needs goes in the brief: the goal, the current state, what to change, what to leave untouched, the project's actual gate commands (discover them from the repo's CLAUDE.md/AGENTS.md/Makefile — do not assume), and a report contract. Tell Grok it will not commit (you will). Keep one task per brief. Full guidance and a template: references/writing-the-brief.md.

2. Dispatch

Send the brief to Grok with the bundled helper. It wraps grok -p, captures the run, and writes a structured result.json — so your only job is "run a command, read a file." (<skill-dir> below is this skill's installed directory — the folder containing this SKILL.md, i.e. the directory you loaded the skill from. Claude Code prints it as "Base directory for this skill" when the skill loads; on other orchestrators use that same directory — if unsure where it landed, run find ~ -name relay.mjs -path '*grok-delegate*' and substitute the directory above it.)

node "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo
# read-only (review/diagnosis; best-effort — verify touchedFiles): add --read-only
# continue the previous Grok session:       add --resume-last  (send only the delta brief)
# hard time limit (watchdog):               add --timeout 2h  (default: off; implementation runs routinely need 1-2h)
# see all options:                          node .../relay.mjs --help

The helper defaults to a write-capable (workspace-write) autonomy profile — --always-approve plus --sandbox workspace — and writes its artifacts to a temp dir, so the repo under review stays clean. It never commits — see step 5. Mechanics, flags, and the result.json shape: references/dispatch-and-poll.md.

3. Wait for completion

The helper blocks until Grok finishes, so back it with whatever your orchestrator offers and resume when it returns:

  • Claude Code: run the Bash call with run_in_background: true; you are notified on completion.
  • Plain shell / other agents: run it in the foreground for short tasks, or background it and poll the result file — … & in bash/zsh (including Git Bash/WSL), or your shell's equivalent (Start-Job in PowerShell, start /b in cmd). The run is done when result.json exists with a status. (A pre-run usage error — bad args or an empty brief — instead exits with code 2 and a stderr message and writes no result file, so check the exit code too. A missing grok binary exits 127 but does write a result.json with status grok_unavailable.)

Do not trust progress trackers over reality: a run is finished when result.json is written and the process has exited. Read the working tree, not a status line. The implementer's full report is the finalMessage field in result.json (also printed in full on stdout between the report markers).

4. Review — do not trust the self-report

Grok's result.json includes its own summary and gate claims. Re-verify, don't accept:

  • Re-run the project's gates yourself (the test/lint/build commands from step 1). Never take "gates passed" on faith.
  • Read the diff against the brief: did Grok do what was asked, nothing more (scope creep) and nothing less? touchedFiles in the result is your starting point.
  • Run the relevant guard skills on the diff if you have them installed (clean-code-guard, test-guard, etc. from guard-skills) — this skill produces the work; those skills judge it.
  • For schema/migration changes, round-trip them; for removals, grep for dangling references.

Full checklist: references/review-and-land.md.

5. Land it

The orchestrator commits. Only after the gates pass and the diff holds:

  • Commit the verified work yourself, with a clear message.
  • If it needs changes, send a delta brief with --resume-last (don't restate the whole task) and review again.

Autonomy model

Grok's default permission mode is ask, which blocks on approval prompts in a headless pipe. The relay therefore always sets autonomy explicitly:

Relay flagWhat Grok getsUse when
(default)--always-approve --sandbox workspaceNormal implementation — writes scoped to the working tree
--read-only--sandbox read-only --always-approveReview / diagnosis — kernel-enforced sandbox, not total (see caveat below)
--full-access--always-approve --sandbox offExplicit opt-in when the task needs unrestricted tools

--always-approve alone would approve all tools (writes, shell, network) — closer to unrestricted than to a workspace-scoped write. Pairing it with --sandbox workspace is what keeps the default safe. Reach for --full-access only when the human asks for it.

--read-only is kernel-enforced, not total. On grok 1.0.25 the read-only sandbox (Seatbelt on macOS, Landlock on Linux) denies grok's own write/search_replace tools and shell redirects with EPERM, so --always-approve only auto-approves tools inside the sandbox. The profile is not total, though: it still permits writes to /tmp, /var/tmp and ~/.grok/, so a repo under one of those paths is not protected, and on macOS it does not restrict child-process network. Always confirm touchedFiles afterward; treat the diff, not the flag, as the guarantee. The relay automates a reporting tripwire: it compares parsed git porcelain and fingerprints the working-tree identity and index entries of Git-visible paths that were already dirty. readOnlyViolation is true when either signal proves a change, false when coverage is complete and detects none, and null when coverage is incomplete. Ignored paths, submodule internals, perfect restores, and attribution of concurrent changes remain outside it, so the diff review stays the guarantee.

Authorization model

Delegation is something the human opts into. Once they have ("run this queue", "proceed"), committing verified, gate-passing work is the agreed contract — that is the whole point. Two limits on that mandate: surface, don't absorb (report Grok's design decisions, defensible-but-unasked turns, and non-blocking nitpicks rather than silently keeping them) and stop for scope changes (if correct completion needs going beyond the brief, ask — don't expand the mandate yourself). The full treatment is in references/review-and-land.md.

References