Skip to content

Goal System Reference

Goal System Reference

Execution Layer Maximum Depth

Overview

The Goal system (pi-goal-list-loop-audit) is the execution layer of the AI Development Pipeline. It takes GitHub issues and automates implementation with independent auditing.

GitHub Issues β†’ [Goal System] β†’ Implemented, Audited, Verified Work

Key Insight: You never manually orchestrate implementation. The Goal system does it for you.


🎯 Three Loop Types

Loop 1: /goal (Single Ordered Goal)

Purpose: Execute one goal with independent verification.

When to Use:

  • One clear task to complete
  • Need semantic judgment (β€œis this done?”)
  • Want audit trail

Commands:

Terminal window
/goal # Drafting: agent grills you
/goal "fix the login bug" # No contract β†’ agent grills first
/goal "Step 1. Step 2. Done when: tests pass." # Has contract β†’ starts now
/goal start "fix the flaky test" # Skip draft, start immediately
/goal status # Show state
/goal pause # Pause
/goal resume # Resume
/goal cancel # Abort
/goal tweak "<new objective>" # Edit in place
/goal archive # View archived goals

Drafting Rules:

  • No args β†’ drafting interview
  • Args without Done when: β†’ agent grills first
  • Args with Done when: β†’ starts immediately
  • /goal start β†’ skip interview

Example:

Terminal window
/goal "Add dark mode toggle. Done when: user can toggle between light/dark themes"

Loop 2: /list (Queue of Goals)

Purpose: Execute multiple goals in sequence.

When to Use:

  • Multiple tasks to complete
  • Bulk import from plan
  • Batch processing

Commands:

Terminal window
/list # Show active + waiting items
/list fix the login bug, add dark mode # Add multiple items
/list plan.md # Import from file
/list <paste checklist> # Multi-line paste
/list next # Skip current, activate next
/list remove <n> # Drop item n
/list clear # Empty the list
/list cancel # Stop the whole list

Key Insight: Order is the default, not the law. /list next <n> picks any item.

Example:

Terminal window
/list fix login bug, add dark mode, write docs, update tests

Loop 3: /loop (Metric-Driven Forever)

Purpose: Continuous improvement until metric plateaus.

When to Use:

  • β€œKeep improving until X”
  • No clear finish line
  • Process that never completes

Commands:

Terminal window
/loop start "reduce TODOs" measure="grep -c TODO src.txt | head -1" direction=min
/loop start "shrink bundle" measure="..." direction=min time=4 tokens=500000
/loop start "keep polishing UI" # Metricless loop
/loop start "reduce TODOs" measure=none max=20 # Metricless with cap
/loop status # Iteration, best, stall
/loop stop # Halt with summary
/loop audit # Project-audit loop

Three Flavors:

  1. Metric loops: Shell command prints honest number
  2. Metricless spec loops: No honest number exists
  3. Audit loops: Fresh audit passes each iteration

Key Insight: No finish line. Runs until you stop it, metric plateaus, or bounds trip.


πŸ›‘οΈ Anti-Bamboozle Architecture

The Problem

If the same AI that writes the code also says β€œI’m done”, how do you know it’s actually done?

The Bamboozle Trap: Agent writes implementation β†’ Agent says β€œI’m done” β†’ Loop trusts them

The Solution: Independent Verification

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ EXECUTOR (writes code) β”‚
β”‚ ───────────────────────────────────────────────────────────── β”‚
β”‚ - Has all skills, extensions, context β”‚
β”‚ - Implements the goal β”‚
β”‚ - Calls complete_goal when done β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ AUDITOR (verifies work) β”‚
β”‚ ───────────────────────────────────────────────────────────── β”‚
β”‚ - Runs in SEPARATE process (pi --mode rpc) β”‚
β”‚ - Has NO extensions, skills, or context β”‚
β”‚ - Only has read/grep/find/ls/bash β”‚
β”‚ - Cannot see the implementation conversation β”‚
β”‚ - Must quote raw evidence for each verification item β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Architectural Guarantees

StageProtection
Goal intakeDrafting + Confirm/Reject dialog; nothing activates unconfirmed
Implementationagent_end-driven continuation loop with 5-minute hard backoff cap
CompletionDetached extension-less auditor process + regression_shield

Regression Shield

The auditor’s report must contain raw tool output for each verification item.

Without shield:

<approved/>

(Auditor says β€œlooks good” - could be rubber-stamp)

With shield:

<approved/>
<evidence>
$ cat src/auth.ts | grep "export function authenticate"
export function authenticate(username: string, password: string): boolean {
$ npm test -- --grep "authentication"
βœ“ should authenticate valid user (12ms)
βœ“ should reject invalid user (8ms)
</evidence>

(Auditor must quote raw output - can’t fake it)

Enforcement: Orchestrator rejects <approved/> without complete evidence β†’ becomes <disapproved/>

Why This Matters

  • Executor can’t fool the auditor
  • Auditor has no context to be biased
  • Evidence is required, not just β€œlooks good”
  • Two independent processes, two perspectives
  • This is unique to Pi - no other AI tool has this

πŸ“‹ Drafting Protocol

How Goals Are Created

  1. User invokes: /goal "do something"
  2. Agent grills: Asks clarifying questions
  3. Agent proposes: propose_goal_draft opens Confirm dialog
  4. User confirms: Goal activates
  5. No goal activates without confirmation

Drafting Interview Example

User: /goal "add user authentication"
Agent: Let me understand what you need:
❓ **Q1** - **Auth Method**: What authentication method?
- OAuth (Google, GitHub)
- Password-based
- Both
➑️ Recommended: Both
❓ **Q2** - **Session Management**: How to manage sessions?
- JWT tokens
- Server-side sessions
- Cookies
➑️ Recommended: JWT
❓ **Q3** - **Security Requirements**: What security level?
- Basic (password hashing)
- Medium (+ rate limiting)
- High (+ 2FA, audit logging)
➑️ Recommended: Medium
[User answers questions]
Agent: Here's the plan:
propose_goal_draft({
objective: "Add user authentication with OAuth and password",
verificationContract: [
"User can register with email/password",
"User can login with OAuth (Google, GitHub)",
"Sessions managed with JWT tokens",
"Rate limiting on login attempts",
"Tests pass"
]
})
[Confirm Dialog appears]
[User confirms]
[Goal activates]

πŸ”„ Status Machine

States

type Status =
| "drafting" // Interview in progress
| "active" // Executing work
| "auditing" // Auditor verifying
| "complete" // Work done, archived
| "paused" // User paused
| "aborted"; // User cancelled

Transitions

drafting β†’ active (user confirms draft)
active β†’ active (continue work)
active β†’ auditing (complete_goal called)
auditing β†’ complete (auditor <approved/>)
auditing β†’ active (auditor <disapproved/>)
active β†’ paused (pause_goal called)
paused β†’ active (user /goal resume)
active β†’ aborted (user /goal cancel)

State Persistence

  • State stored in .pi-glla/active.jsonl
  • Each line is a state transition
  • Deterministic compaction from JSONL
  • Protects against model-generated summaries losing fidelity

🚨 Recovery Systems

Stall Detection

Self-watchdog (15s heartbeat):

  • Active goal + idle session + nothing scheduled + 60s quiet β†’ re-fire
  • Three consecutive no-tool turns β†’ pause
  • No external watchdog plugin needed

Wedge Alert (30 minutes):

  • Busy + no activity for 30 minutes β†’ warning + notify
  • Tune in /glla settings (0 = off)

Quota Walls

When provider hits limits:

glla: ⟦⏳ QUOTA WALL · next probe in 10m 48s⟧ · 1 queued
β”œβ”€ QUOTA WALL Β· Token Plan usage limit Β· 1 waiting in list
β”œβ”€ waiting β€” nothing for you to do Β· next probe in 10m 48s

Recovery envelope: 15m β†’ 30m β†’ 1h β†’ 2h β†’ 4h β†’ 5h (cap 5h, automatic window 24h)

Knowledge-window escalation: Quota/billing/auth failures escalate faster (3 minutes instead of 15)

Session Handoff

Recovery crosses pi’s lifecycle:

  1. session_shutdown β†’ persists continuation debt
  2. Fresh session_start β†’ consumes debt, continues
  3. Stale handles refuse mutations (fail closed)
  4. User quit = not implicit resume consent

Orphan handling:

  • Dirty stacked states β†’ most recent activity keeps slot
  • Loser archived (recoverable)
  • No picker, no arbitration chores

User Aborts

Aborted turn:

  • Exempt from stall accounting
  • Stands chain down (no auto re-fire)
  • Stand-down survives heartbeat
  • 5 consecutive aborts β†’ loud pause

πŸ“Š Status Display

Live TUI

Persistent status segment shows current state:

glla: [β–β–‚β–„β–†β–ˆβ–† LIVE Β· WORKING] 1m 09s Β· last stream 11s ago Β· 3 queued
glla: [QUEUED] 44s Β· 18 queued

Activity Indicators

IndicatorMeaning
LIVE Β· WORKINGFresh stream/tool activity arriving
BUSYpi occupied, no fresh stream evidence
QUEUEDContinuation waiting to start
IDLEActive item, no recent work
auditor …Detached verifier queued/running
QUOTA WALLProvider rejected request

Queue Trail

For long-running /list work:

● Fix the current issue Β· list item Β· active Β· 42m
β”œβ”€ βœ“ bash tests/display.test.ts (35s) Β· next: update docs
β”œβ”€ ↳ 23 waiting Β· up next: refresh the release notes Β· waiting 12m 04s
└─ 23 queued Β· /list Β· /glla

βš™οΈ Configuration

/glla Settings

Open /glla to edit:

  • Auditor model and thinking level
  • Auditor fallback model
  • Notify command and settings
  • Auto-resume, auto-accept drafts
  • Main session backups
  • Recovery cadence
  • Audit cap/report size

Resolution Order

Project > Global > Defaults

Exception: autoResume is global-only (per-project opt-ins silently overrode global hold)

Main Model Fallbacks

Ordered list of backup models:

mainModelFallbacks: [
"anthropic/claude-sonnet-4-20250514",
"openai/gpt-4o",
"google/gemini-2.0-flash"
]

Provider error β†’ rotate through authenticated candidates β†’ all fail β†’ park and probe


πŸ”§ Integration with Pipeline

Input: GitHub Issues

Goals are created from GitHub issues:

/goal "Fix login bug. Done when: tests pass"

The issue provides:

  • Objective (what to do)
  • Verification contract (how to know it’s done)
  • Context (background information)

Output: Verified Work

Goals output:

  • Implemented code
  • Audit trail (evidence blocks)
  • Verification status (approved/disapproved)
  • Archived goal for history

Execution Skills

Within goals, use:

  • Impeccable: UI/design work
  • Ponytail: Minimal solutions
  • Pi Agent skills: Platform-specific work

πŸ“ Files

.pi-glla/
β”œβ”€β”€ active.jsonl # Current state
β”œβ”€β”€ session-handoff.json # Recovery data
β”œβ”€β”€ archive/ # Completed goals
β”œβ”€β”€ audit-jobs/ # Auditor work
└── continuation-dispatch.json # Dispatch tracking

πŸŽ“ Best Practices

  1. Always draft first: Let the agent grill you
  2. Write clear contracts: β€œDone when: X, Y, Z”
  3. Use /list for batches: Multiple related tasks
  4. Use /loop for continuous improvement: No clear finish line
  5. Trust the auditor: Two perspectives are better than one
  6. Check /goal status: Know what’s happening
  7. Use /glla for configuration: Tune to your workflow