From 3fed6f3cfe293c05ffb456c6dc50ee630b1ac5eb Mon Sep 17 00:00:00 2001 From: Leone Martins Date: Fri, 31 Jul 2026 20:53:56 -0300 Subject: [PATCH 1/2] feat(commands): add Antigravity CLI (agy) as santa-loop Reviewer B option MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Santa-loop's Reviewer B only tried codex and gemini before falling back to a same-family Claude reviewer, losing model diversity when neither was installed. Detect agy (Antigravity CLI, ~/.local/bin/agy) as a third option, using gemini-3.6-flash-high — despite the "flash" name, it currently outranks gemini-3.1-pro-high on every published coding/agentic benchmark (SWE-Bench Pro, Terminal-Bench, MLE-Bench), with Pro only ahead on PhD-level reasoning benchmarks that don't apply to code review. Also documents never pointing agy at a Claude model, which would collapse Reviewer A/B model diversity entirely. --- commands/santa-loop.md | 19 ++++++++++++++----- 1 file changed, 14 insertions(+), 5 deletions(-) diff --git a/commands/santa-loop.md b/commands/santa-loop.md index 111087966..015d9ea24 100644 --- a/commands/santa-loop.md +++ b/commands/santa-loop.md @@ -76,6 +76,7 @@ First, detect which CLIs are available: ```bash command -v codex >/dev/null 2>&1 && echo "codex" || true command -v gemini >/dev/null 2>&1 && echo "gemini" || true +command -v agy >/dev/null 2>&1 && echo "agy" || true ``` Build the reviewer prompt (identical rubric + instructions as Reviewer A) and write it to a unique temp file: @@ -86,7 +87,7 @@ cat > "$PROMPT_FILE" << 'EOF' EOF ``` -Use the first available CLI: +Use the first available CLI, in this order: **Codex CLI** (if installed) ```bash @@ -100,7 +101,14 @@ gemini -p "$(cat "$PROMPT_FILE")" -m gemini-2.5-pro rm -f "$PROMPT_FILE" ``` -**Claude Agent fallback** (only if neither `codex` nor `gemini` is installed) +**Antigravity CLI** (if installed and neither codex nor gemini is) +```bash +agy -p "$(cat "$PROMPT_FILE")" --model gemini-3.6-flash-high --sandbox +rm -f "$PROMPT_FILE" +``` +Despite the "flash" name, this outranks `gemini-3.1-pro-high` on every published coding/agentic benchmark (SWE-Bench Pro, Terminal-Bench, MLE-Bench) — Pro only leads on PhD-level reasoning benchmarks (GPQA, HLE), which aren't relevant to code review. Don't "correct" this back to a `-pro-` model by name alone; check current benchmarks first, since generation-over-tier ordering shifts release to release. Run `agy models` to see the current catalog before assuming this is stale. + +**Claude Agent fallback** (only if none of `codex`, `gemini`, or `agy` is installed) Launch a second Claude Agent (subagent_type: `code-reviewer`, model: `opus`). Log a warning that both reviewers share the same model family — true model diversity was not achieved but context isolation is still enforced. In all cases, the reviewer must return the same structured JSON verdict as Reviewer A. @@ -166,9 +174,10 @@ Result: [PUSHED / ESCALATED TO USER] ## Notes - Reviewer A (Claude Opus) always runs — guarantees at least one strong reviewer regardless of tooling. -- Model diversity is the goal for Reviewer B. GPT-5.4 or Gemini 2.5 Pro gives true independence — different training data, different biases, different blind spots. The Claude-only fallback still provides value via context isolation but loses model diversity. -- Strongest available models are used: Opus for Reviewer A, GPT-5.4 or Gemini 2.5 Pro for Reviewer B. -- External reviewers run with `--sandbox read-only` (Codex) to prevent repo mutation during review. +- Model diversity is the goal for Reviewer B. GPT-5.4, Gemini 2.5 Pro, or Antigravity's Gemini 3.6 Flash (via `agy`) all give true independence — different training data, different biases, different blind spots. The Claude-only fallback still provides value via context isolation but loses model diversity. +- Strongest available models are used: Opus for Reviewer A, GPT-5.4, Gemini 2.5 Pro, or Gemini 3.6 Flash High (`agy`) for Reviewer B, in that priority order. +- Never point `agy` at a Claude model (`claude-sonnet-4-6`, `claude-opus-4-6-thinking`) — Reviewer A is already Claude Opus, so that would eliminate model diversity entirely. +- External reviewers run with `--sandbox read-only` (Codex) or `--sandbox` (`agy`) to prevent repo mutation during review. - Fresh reviewers each round prevents anchoring bias from prior findings. - The rubric is the most important input. Tighten it if reviewers rubber-stamp or flag subjective style issues. - Commits happen on NAUGHTY rounds so fixes are preserved even if the loop is interrupted. From d73009bbd17ece34b193c107662c3d4ffe2ffee7 Mon Sep 17 00:00:00 2001 From: haelyra <49814733+haelyra@users.noreply.github.com> Date: Tue, 11 Aug 2026 13:39:42 -0400 Subject: [PATCH 2/2] fix(commands): route Santa reviewer through Antigravity wrapper --- commands/santa-loop.md | 21 ++++++++++++--------- 1 file changed, 12 insertions(+), 9 deletions(-) diff --git a/commands/santa-loop.md b/commands/santa-loop.md index 015d9ea24..29a3a6031 100644 --- a/commands/santa-loop.md +++ b/commands/santa-loop.md @@ -76,7 +76,7 @@ First, detect which CLIs are available: ```bash command -v codex >/dev/null 2>&1 && echo "codex" || true command -v gemini >/dev/null 2>&1 && echo "gemini" || true -command -v agy >/dev/null 2>&1 && echo "agy" || true +test -x "$HOME/.claude/bin/codeagent-wrapper" && echo "antigravity" || true ``` Build the reviewer prompt (identical rubric + instructions as Reviewer A) and write it to a unique temp file: @@ -101,14 +101,18 @@ gemini -p "$(cat "$PROMPT_FILE")" -m gemini-2.5-pro rm -f "$PROMPT_FILE" ``` -**Antigravity CLI** (if installed and neither codex nor gemini is) +**Antigravity backend** (if the CCG wrapper is installed and neither Codex nor Gemini is) ```bash -agy -p "$(cat "$PROMPT_FILE")" --model gemini-3.6-flash-high --sandbox +REVIEWER_ROLE="$HOME/.claude/.ccg/prompts/antigravity/reviewer.md" +{ + printf 'ROLE_FILE: %s\n' "$REVIEWER_ROLE" + cat "$PROMPT_FILE" +} | "$HOME/.claude/bin/codeagent-wrapper" --backend antigravity - "$PWD" rm -f "$PROMPT_FILE" ``` -Despite the "flash" name, this outranks `gemini-3.1-pro-high` on every published coding/agentic benchmark (SWE-Bench Pro, Terminal-Bench, MLE-Bench) — Pro only leads on PhD-level reasoning benchmarks (GPQA, HLE), which aren't relevant to code review. Don't "correct" this back to a `-pro-` model by name alone; check current benchmarks first, since generation-over-tier ordering shifts release to release. Run `agy models` to see the current catalog before assuming this is stale. +Do not hardcode a model ID here. The wrapper owns Antigravity model selection, so the workflow stays compatible as the backend catalog changes. -**Claude Agent fallback** (only if none of `codex`, `gemini`, or `agy` is installed) +**Claude Agent fallback** (only if Codex, Gemini, and the Antigravity wrapper are unavailable) Launch a second Claude Agent (subagent_type: `code-reviewer`, model: `opus`). Log a warning that both reviewers share the same model family — true model diversity was not achieved but context isolation is still enforced. In all cases, the reviewer must return the same structured JSON verdict as Reviewer A. @@ -174,10 +178,9 @@ Result: [PUSHED / ESCALATED TO USER] ## Notes - Reviewer A (Claude Opus) always runs — guarantees at least one strong reviewer regardless of tooling. -- Model diversity is the goal for Reviewer B. GPT-5.4, Gemini 2.5 Pro, or Antigravity's Gemini 3.6 Flash (via `agy`) all give true independence — different training data, different biases, different blind spots. The Claude-only fallback still provides value via context isolation but loses model diversity. -- Strongest available models are used: Opus for Reviewer A, GPT-5.4, Gemini 2.5 Pro, or Gemini 3.6 Flash High (`agy`) for Reviewer B, in that priority order. -- Never point `agy` at a Claude model (`claude-sonnet-4-6`, `claude-opus-4-6-thinking`) — Reviewer A is already Claude Opus, so that would eliminate model diversity entirely. -- External reviewers run with `--sandbox read-only` (Codex) or `--sandbox` (`agy`) to prevent repo mutation during review. +- Model diversity is the goal for Reviewer B. Codex, Gemini, or the Antigravity backend provides a different provider family from Reviewer A. The Claude-only fallback still provides value via context isolation but loses model diversity. +- Use each backend's maintained model-selection contract. Do not pin a transient Antigravity model ID in this workflow. +- External reviewers use their supported restricted execution path. Codex runs with `--sandbox read-only`; Antigravity runs through the CCG wrapper instead of an undocumented direct CLI contract. - Fresh reviewers each round prevents anchoring bias from prior findings. - The rubric is the most important input. Tighten it if reviewers rubber-stamp or flag subjective style issues. - Commits happen on NAUGHTY rounds so fixes are preserved even if the loop is interrupted.