Goal System Reference
Goal System Reference
Execution Layer Maximum DepthOverview
The Goal system (pi-goal-list-loop-audit) is the execution layer of the AI Development Pipeline. It takes GitHub issues and automates implementation with independent auditing.
GitHub Issues β [Goal System] β Implemented, Audited, Verified WorkKey Insight: You never manually orchestrate implementation. The Goal system does it for you.
π― Three Loop Types
Loop 1: /goal (Single Ordered Goal)
Purpose: Execute one goal with independent verification.
When to Use:
- One clear task to complete
- Need semantic judgment (βis this done?β)
- Want audit trail
Commands:
/goal # Drafting: agent grills you/goal "fix the login bug" # No contract β agent grills first/goal "Step 1. Step 2. Done when: tests pass." # Has contract β starts now/goal start "fix the flaky test" # Skip draft, start immediately/goal status # Show state/goal pause # Pause/goal resume # Resume/goal cancel # Abort/goal tweak "<new objective>" # Edit in place/goal archive # View archived goalsDrafting Rules:
- No args β drafting interview
- Args without
Done when:β agent grills first - Args with
Done when:β starts immediately /goal startβ skip interview
Example:
/goal "Add dark mode toggle. Done when: user can toggle between light/dark themes"Loop 2: /list (Queue of Goals)
Purpose: Execute multiple goals in sequence.
When to Use:
- Multiple tasks to complete
- Bulk import from plan
- Batch processing
Commands:
/list # Show active + waiting items/list fix the login bug, add dark mode # Add multiple items/list plan.md # Import from file/list <paste checklist> # Multi-line paste/list next # Skip current, activate next/list remove <n> # Drop item n/list clear # Empty the list/list cancel # Stop the whole listKey Insight: Order is the default, not the law. /list next <n> picks any item.
Example:
/list fix login bug, add dark mode, write docs, update testsLoop 3: /loop (Metric-Driven Forever)
Purpose: Continuous improvement until metric plateaus.
When to Use:
- βKeep improving until Xβ
- No clear finish line
- Process that never completes
Commands:
/loop start "reduce TODOs" measure="grep -c TODO src.txt | head -1" direction=min/loop start "shrink bundle" measure="..." direction=min time=4 tokens=500000/loop start "keep polishing UI" # Metricless loop/loop start "reduce TODOs" measure=none max=20 # Metricless with cap/loop status # Iteration, best, stall/loop stop # Halt with summary/loop audit # Project-audit loopThree Flavors:
- Metric loops: Shell command prints honest number
- Metricless spec loops: No honest number exists
- Audit loops: Fresh audit passes each iteration
Key Insight: No finish line. Runs until you stop it, metric plateaus, or bounds trip.
π‘οΈ Anti-Bamboozle Architecture
The Problem
If the same AI that writes the code also says βIβm doneβ, how do you know itβs actually done?
The Bamboozle Trap: Agent writes implementation β Agent says βIβm doneβ β Loop trusts them
The Solution: Independent Verification
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ EXECUTOR (writes code) ββ βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ ββ - Has all skills, extensions, context ββ - Implements the goal ββ - Calls complete_goal when done ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ AUDITOR (verifies work) ββ βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ ββ - Runs in SEPARATE process (pi --mode rpc) ββ - Has NO extensions, skills, or context ββ - Only has read/grep/find/ls/bash ββ - Cannot see the implementation conversation ββ - Must quote raw evidence for each verification item ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββArchitectural Guarantees
| Stage | Protection |
|---|---|
| Goal intake | Drafting + Confirm/Reject dialog; nothing activates unconfirmed |
| Implementation | agent_end-driven continuation loop with 5-minute hard backoff cap |
| Completion | Detached extension-less auditor process + regression_shield |
Regression Shield
The auditorβs report must contain raw tool output for each verification item.
Without shield:
<approved/>(Auditor says βlooks goodβ - could be rubber-stamp)
With shield:
<approved/><evidence>$ cat src/auth.ts | grep "export function authenticate"export function authenticate(username: string, password: string): boolean {$ npm test -- --grep "authentication" β should authenticate valid user (12ms) β should reject invalid user (8ms)</evidence>(Auditor must quote raw output - canβt fake it)
Enforcement: Orchestrator rejects <approved/> without complete evidence β becomes <disapproved/>
Why This Matters
- Executor canβt fool the auditor
- Auditor has no context to be biased
- Evidence is required, not just βlooks goodβ
- Two independent processes, two perspectives
- This is unique to Pi - no other AI tool has this
π Drafting Protocol
How Goals Are Created
- User invokes:
/goal "do something" - Agent grills: Asks clarifying questions
- Agent proposes:
propose_goal_draftopens Confirm dialog - User confirms: Goal activates
- No goal activates without confirmation
Drafting Interview Example
User: /goal "add user authentication"
Agent: Let me understand what you need:
β **Q1** - **Auth Method**: What authentication method? - OAuth (Google, GitHub) - Password-based - Both
β‘οΈ Recommended: Both
β **Q2** - **Session Management**: How to manage sessions? - JWT tokens - Server-side sessions - Cookies
β‘οΈ Recommended: JWT
β **Q3** - **Security Requirements**: What security level? - Basic (password hashing) - Medium (+ rate limiting) - High (+ 2FA, audit logging)
β‘οΈ Recommended: Medium
[User answers questions]
Agent: Here's the plan:
propose_goal_draft({ objective: "Add user authentication with OAuth and password", verificationContract: [ "User can register with email/password", "User can login with OAuth (Google, GitHub)", "Sessions managed with JWT tokens", "Rate limiting on login attempts", "Tests pass" ]})
[Confirm Dialog appears][User confirms][Goal activates]π Status Machine
States
type Status = | "drafting" // Interview in progress | "active" // Executing work | "auditing" // Auditor verifying | "complete" // Work done, archived | "paused" // User paused | "aborted"; // User cancelledTransitions
drafting β active (user confirms draft)active β active (continue work)active β auditing (complete_goal called)auditing β complete (auditor <approved/>)auditing β active (auditor <disapproved/>)active β paused (pause_goal called)paused β active (user /goal resume)active β aborted (user /goal cancel)State Persistence
- State stored in
.pi-glla/active.jsonl - Each line is a state transition
- Deterministic compaction from JSONL
- Protects against model-generated summaries losing fidelity
π¨ Recovery Systems
Stall Detection
Self-watchdog (15s heartbeat):
- Active goal + idle session + nothing scheduled + 60s quiet β re-fire
- Three consecutive no-tool turns β pause
- No external watchdog plugin needed
Wedge Alert (30 minutes):
- Busy + no activity for 30 minutes β warning + notify
- Tune in
/gllasettings (0 = off)
Quota Walls
When provider hits limits:
glla: β¦β³ QUOTA WALL Β· next probe in 10m 48sβ§ Β· 1 queuedββ QUOTA WALL Β· Token Plan usage limit Β· 1 waiting in listββ waiting β nothing for you to do Β· next probe in 10m 48sRecovery envelope: 15m β 30m β 1h β 2h β 4h β 5h (cap 5h, automatic window 24h)
Knowledge-window escalation: Quota/billing/auth failures escalate faster (3 minutes instead of 15)
Session Handoff
Recovery crosses piβs lifecycle:
session_shutdownβ persists continuation debt- Fresh
session_startβ consumes debt, continues - Stale handles refuse mutations (fail closed)
- User quit = not implicit resume consent
Orphan handling:
- Dirty stacked states β most recent activity keeps slot
- Loser archived (recoverable)
- No picker, no arbitration chores
User Aborts
Aborted turn:
- Exempt from stall accounting
- Stands chain down (no auto re-fire)
- Stand-down survives heartbeat
- 5 consecutive aborts β loud pause
π Status Display
Live TUI
Persistent status segment shows current state:
glla: [ββββββ LIVE Β· WORKING] 1m 09s Β· last stream 11s ago Β· 3 queuedglla: [QUEUED] 44s Β· 18 queuedActivity Indicators
| Indicator | Meaning |
|---|---|
LIVE Β· WORKING | Fresh stream/tool activity arriving |
BUSY | pi occupied, no fresh stream evidence |
QUEUED | Continuation waiting to start |
IDLE | Active item, no recent work |
auditor β¦ | Detached verifier queued/running |
QUOTA WALL | Provider rejected request |
Queue Trail
For long-running /list work:
β Fix the current issue Β· list item Β· active Β· 42mββ β bash tests/display.test.ts (35s) Β· next: update docsββ β³ 23 waiting Β· up next: refresh the release notes Β· waiting 12m 04sββ 23 queued Β· /list Β· /gllaβοΈ Configuration
/glla Settings
Open /glla to edit:
- Auditor model and thinking level
- Auditor fallback model
- Notify command and settings
- Auto-resume, auto-accept drafts
- Main session backups
- Recovery cadence
- Audit cap/report size
Resolution Order
Project > Global > Defaults
Exception: autoResume is global-only (per-project opt-ins silently overrode global hold)
Main Model Fallbacks
Ordered list of backup models:
mainModelFallbacks: [ "anthropic/claude-sonnet-4-20250514", "openai/gpt-4o", "google/gemini-2.0-flash"]Provider error β rotate through authenticated candidates β all fail β park and probe
π§ Integration with Pipeline
Input: GitHub Issues
Goals are created from GitHub issues:
/goal "Fix login bug. Done when: tests pass"The issue provides:
- Objective (what to do)
- Verification contract (how to know itβs done)
- Context (background information)
Output: Verified Work
Goals output:
- Implemented code
- Audit trail (evidence blocks)
- Verification status (approved/disapproved)
- Archived goal for history
Execution Skills
Within goals, use:
- Impeccable: UI/design work
- Ponytail: Minimal solutions
- Pi Agent skills: Platform-specific work
π Files
.pi-glla/βββ active.jsonl # Current stateβββ session-handoff.json # Recovery dataβββ archive/ # Completed goalsβββ audit-jobs/ # Auditor workβββ continuation-dispatch.json # Dispatch trackingπ Best Practices
- Always draft first: Let the agent grill you
- Write clear contracts: βDone when: X, Y, Zβ
- Use /list for batches: Multiple related tasks
- Use /loop for continuous improvement: No clear finish line
- Trust the auditor: Two perspectives are better than one
- Check
/goal status: Know whatβs happening - Use
/gllafor configuration: Tune to your workflow